Data Labeling Startups funded by Y Combinator (YC) 2026

September 2026

Browse 20 of the top Data Labeling startups funded by Y Combinator.

We also have a Startup Directory where you can search through over 5,000 companies.

  • Moving Atoms
    Moving Atoms
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Our first model, Atom 1, pre-trained on the internets video, and fine tuned for physics. It has been action conditioned, so you can simulate robot actions within it. This lets Robot Labs train their robots using GPUs instead of relying on human collected manual data that is slow and expensive. In unseen and hard to replicate environments, this is the only path to getting ANY training data.
    data-labeling
    robotics
    virtual-reality
    artificial-intelligence
    hard-tech
  • CoArena
    CoArena
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    CoArena is a live arena where anyone can use the world's top computer-use models racing two of them on the same real computer task and judging which one did it better. Every battle becomes something the AI labs can't build themselves, an honest test of their agents on real work, and the data to make them better.
    aiops
    reinforcement-learning
    data-labeling
    artificial-intelligence
    machine-learning
  • Markov
    Markov
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Markov builds expert computer-use datasets for frontier AI labs. Our datasets include screen recordings + mouse/keyboard actions of professionals using CAD, design and enterprise software. We work with the top AI labs and we’ve gotten more than 350k downloads on HuggingFace. Our long-term goal is to be the full stack platform for computer-use AI, starting with data and then the infrastructure to train and deploy computer-use models. Dev and Harish are technical co-founders. Dev studied aerospace engineering at IIT Madras and was the youngest team member at Sarvam - India’s top AI lab. Harish is an Emergent Ventures grantee and his prior robotics work was featured in top national publications.
    reinforcement-learning
    data-labeling
    data-engineering
  • Shofo
    Shofo
    Y Combinator LogoW2026
    Active • 4 employees • San Francisco
    We are building the world’s largest video library. We've aggregated billions of videos into a searchable index and use agents to find and label the exact datasets a lab needs on demand. If a lab needs 100K hours of cooking videos where someone is holding a pan, with reasoning annotations on top, our agents search the index, extract the matching subset, route it through our labeling pipeline, and deliver a custom dataset in days, not months.
    data-labeling
    infrastructure
    artificial-intelligence
  • Datoric
    Datoric
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Datoric develops custom datasets for voice models, robotics, and world models, treating research, collection, verification, and production as one continuous process. We work closely with frontier model teams to turn emerging limitations and model failures into testable data hypotheses, while running our own experiments ahead of customer demand. This allows us to operationalize validated methods into repeatable collection systems at scale. Data is collected through private invite-only applications separated by modality, customer, and trust level. Every submission remains linked to the contributor, device, task, session, consent, rights, and processing history that produced it, giving our internal QA and fraud models the context to detect problems that may appear legitimate in the finished file. Each collection reveals new failure cases and quality signals that improve the systems behind the next dataset. Once a collection method is validated, we scale it up as a solution to the model failure and it also becomes a reusable data recipe for future custom projects or independently collected, rights-cleared data products.
    data-labeling
    machine-learning
    robotics
    artificial-intelligence
  • Sciloop
    Sciloop
    Y Combinator LogoF2025
    Active • 6 employees • San Francisco
    Sciloop creates expert-level math and physics problems that frontier AI models can't solve, then sells the data to AI labs for training and evaluation. Our problems are created by IPhO and IMO medalists — the top 0.01% of STEM talent globally. On our benchmark, models like GPT 5.4 Pro and Gemini 3.1 Pro score 0-5% on our hardest problems. We work with AI labs to supply continuous, fresh training data that pushes the frontier of mathematical and scientific reasoning. Founded by Bilal and Osman, International Physics Olympiad medalists from MIT with hands-on ML research experience at MIT CSAIL.
    ai
    big-data
    data-labeling
    marketplace
  • Liva AI
    Liva AI
    Y Combinator LogoS2025
    Active • 6 employees • San Francisco
    Helping to build more socially intelligent AI, starting with voice.
    b2b
    big-data
    data-labeling
    marketplace
    artificial-intelligence
  • Besimple AI
    Besimple AI
    Y Combinator LogoP2025
    Active • 10 employees • San Francisco
    We are building the data layer for AI, starting with audio. We start with data collection, curating our own proprietary set of diverse conversational data covering a wide range of languages, dialects and accents. We then leverage human expert audio annotators and our own annotation platform to process audio data for Automatic Speech Recognition. With human level transcription and diarization, our data help push the audio model frontier. Today we have over millions of hours of conversational data, and growing. If you need audio data for training or evaluating your voice models or voice agents, reach out! We offer flexible licensing deals that work for startups and enterprises, with minimal process. Audio data should besimple :)
    data-labeling
    aiops
    ai
  • Cartpole
    Cartpole
    Y Combinator LogoP2025
    Active • 2 employees • San Francisco
    We're creating reinforcement learning environments for training frontier models.
    reinforcement-learning
    ml
    artificial-intelligence
    data-labeling
  • Sureform
    Sureform
    Y Combinator LogoP2025
    Active • 2 employees • San Francisco
    We build post-training datasets and RL environments grounded in real enterprise workflows to advance frontier agents.
    data-labeling
    robotics
    marketplace
    ai
  • AfterQuery
    AfterQuery
    Y Combinator LogoW2025
    Active • 30 employees • San Francisco
    AfterQuery is an applied research lab curating data solutions for frontier foundation model development. Serving every frontier AI lab.
    b2b
    artificial-intelligence
    ai
    big-data
    data-labeling
  • Unbound
    Unbound
    Y Combinator LogoS2024
    Active • 7 employees • San Francisco
    cybersecurity
    artificial-intelligence
    privacy
    data-labeling
  • Sieve
    Sieve
    Y Combinator LogoW2022
    Active • 34 employees • San Francisco
    Sieve builds the data and environments frontier AI labs use to train the next generation of multimodal systems. AI is moving beyond chatbots into video, audio, images, software, robotics, and interactive worlds. The next generation of models will need to understand how the world looks, sounds, moves, responds, and changes over time. Progress is bottlenecked by one thing: high-quality data. Sieve brings together exabyte-scale infrastructure, novel multimodal understanding techniques, large-scale sourcing, and deep research partnerships to create datasets and environments with unmatched precision, quality, and speed. This has earned the trust of frontier AI labs, Fortune 100 companies, and fast-growing AI startups working on generative media, robotics, computer use, world models, and agentic systems.
    video
    developer-tools
    data-engineering
    data-labeling
    ai
  • Spade
    Spade
    Y Combinator LogoW2022
    Active • 30 employees • New York City
    We take messy transaction data and turn it into structured, verified records — with a platform, models, and tooling that help our customers use it everywhere it matters.
    fintech
    machine-learning
    payments
    data-labeling
    artificial-intelligence
  • Lightly
    Lightly
    Y Combinator LogoS2021
    Active • 5 employees • Zürich, Switzerland
    When ML teams send their data to companies like Scale.ai for labeling, most can only afford to label 1% or less of their datasets. But today they don’t have a good way to pick which 1% to label. We help them pick the best 1% of their data to label. By labeling the most representative data, they significantly improve model accuracy at the same cost.
    machine-learning
    data-labeling
  • Centaur
    Centaur
    Y Combinator LogoW2019
    Active • 45 employees • Boston
    The best AI models aren’t just trained and evaluated with human data; they’re built with superhuman data. The strongest datasets emerge through collective intelligence, where humans and machines work together to outperform either one alone. At Centaur, we create superior quality data by turning annotation into an arena where experts and AI compete.
    data-labeling
    crowdsourcing
    data-science
    artificial-intelligence
  • Sepal AI
    Sepal AI
    Y Combinator LogoS2024
    Acquired • 15 employees • San Francisco
    Sepal is a data research company on a mission to advance human knowledge and capabilities through safe AI. We partner with the world’s leading AI labs and enterprises to help their models get better at the tasks people actually want them to do. We’ve built a Cloud-Native Agent Dataset Factory which turns the process of generating evaluation and training data from manual, inconsistent, and labor-intensive into something automated, standardized, and scalable. Sepal AI was founded in 2024 by engineers and operators from Vercel and Turing. We went through Y Combinator, raised several million dollars from leading investors, and already count multiple Fortune 500s and top AI research labs as paying customers.
    data-labeling
    aiops
    reinforcement-learning
    ai
  • Deasy Labs
    Deasy Labs
    Y Combinator LogoS2023
    Acquired • 8 employees • New York City
    Deasy Labs was acquired by Collibra in July 2025 (global leader in enterprise data governance). Deasy Labs provides metadata orchestration for AI workflows. Deasie's platform provides the best way for AI teams to create and embed high-quality, customized metadata into their AI workflows (e.g., RAG, Agentic frameworks). Our three founders (from Amazon, McKinsey/QuantumBlack & MIT) previously built an ML data governance tool from 0 to 1 within McKinsey, which we deployed with 11 Fortune 500 companies. We saw in early 2023 the ability to create high-quality metadata (without reliance on domain experts) would be a key factor in achieving the accuracy & speed in GenAI applications required for production. Our investors include General Catalyst, Y Combinator, RTP Global and world experts in enterprise data. Website: https://deasylabs.com
    ai-assistant
    data-labeling
    databases
    big-data
    artificial-intelligence
  • JumpWire
    JumpWire
    Y Combinator LogoW2022
    Acquired • 2 employees • New York City
    JumpWire is a data protection platform that adds advanced data security controls between APIs, applications and databases. JumpWire automatically identifies sensitive properties inside large data sets and gives developers full control over which people and applications can access or update records containing sensitive info. Examples uses include restricting who can read customer PII to members of the customer service team, giving on-call engineers elevated access to production, or splitting user records between regions for GDPR purposes. JumpWire’s approach to securing data in-place minimizes the risk of data leaks exposing sensitive information or mishandling by other applications and vendors. The exact security scheme applied to data is defined by policies that align with an organization’s existing InfoSec program. JumpWire helps companies who maintain information security with compliance programs such as SOC or HIPAA. They are processing sensitive data, often from their own customers, and exceed security best practices as a competitive advantage. JumpWire provides defense at depth to data and sits alongside access controls and Layer 4 encryption to provide a comprehensive data security solution. JumpWire is unique from solutions such as data vaults by installing inside our customers’ own infrastructure and clouds. It is interoperable with existing applications and databases, which eliminates the need for large data migrations or code refactoring. Lower-level approaches to data security, such as encryption at rest, are too blunt and lack the ability to differentiate between properties in the data itself. Its scope is limited to physical storage, and security is lost as soon as an application or query loads the data.
    security
    data-labeling
    databases